assembled reads Search Results


90
Oxford Nanopore oxford nanopore long read assemblies
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Oxford Nanopore Long Read Assemblies, supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/bio_rxiv__2024__11__01__618662-39-2-14?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
oxford nanopore long read assemblies - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Oxford Nanopore long sequencing read-based assembly
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Long Sequencing Read Based Assembly, supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc11436853-229-2-13?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
long sequencing read-based assembly - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Broad Institute Inc long read-based assembly
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Long Read Based Assembly, supplied by Broad Institute Inc, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/10__1094_slash_mpmi___09___18___0265___ta-765-13-25?v=Broad+Institute+Inc
Average 90 stars, based on 1 article reviews
long read-based assembly - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Celera assembler fork for long reads
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Assembler Fork For Long Reads, supplied by Celera, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc06808082-93-9-8?v=Celera
Average 90 stars, based on 1 article reviews
assembler fork for long reads - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Celera maryland super-read celera assembler masurca v3.2.3
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Maryland Super Read Celera Assembler Masurca V3.2.3, supplied by Celera, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc07250882__41467_2020_16284_MOESM1_ESM-35-8-5?v=Celera
Average 90 stars, based on 1 article reviews
maryland super-read celera assembler masurca v3.2.3 - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Oxford Nanopore assemble reads
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Assemble Reads, supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc11317156-192-35-47?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
assemble reads - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
KAUST Core Labs assembly read error correction tool karect
(a) Read based detection of Patescibacterium . A total of 580 SRA <t>metagenomes,</t> out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .
Assembly Read Error Correction Tool Karect, supplied by KAUST Core Labs, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc07708060-59-19-18?v=KAUST+Core+Labs
Average 90 stars, based on 1 article reviews
assembly read error correction tool karect - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Oxford Nanopore contigs assembled from oxford nanopore minion long-reads
Descriptive characteristics for three draft taro genome assemblies. The pseudochromosome-level taro assembly (“Ps_chr”) was composed from a linked-read assembly (“LR”) that was gap-filled using <t> contigs </t> assembled from nanopore MinION long-reads (merge step, “Merged”), filtered for assembly artifacts, and then concatenated into pseudochromosomes using a linkage map. Kilobase = kb
Contigs Assembled From Oxford Nanopore Minion Long Reads, supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc07407455-212-22-25?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
contigs assembled from oxford nanopore minion long-reads - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Oxford Nanopore long read sequencing technology used for de novo genome assembly, isoform identification, and detecting epigenetic modifications.
Evaluating the pros and cons of selected sequencing and mapping technologies.
Long Read Sequencing Technology Used For De Novo Genome Assembly, Isoform Identification, And Detecting Epigenetic Modifications., supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc08041138-1-18-0?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
long read sequencing technology used for de novo genome assembly, isoform identification, and detecting epigenetic modifications. - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
SourceForge net tool for de novo assembly of short reads with robust error detection
Evaluating the pros and cons of selected sequencing and mapping technologies.
Tool For De Novo Assembly Of Short Reads With Robust Error Detection, supplied by SourceForge net, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pm19679362-101-0-1?v=SourceForge+net
Average 90 stars, based on 1 article reviews
tool for de novo assembly of short reads with robust error detection - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
PopulationGenetics trinity short read assembler
Evaluating the pros and cons of selected sequencing and mapping technologies.
Trinity Short Read Assembler, supplied by PopulationGenetics, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/pmc05068948-101-6-53?v=PopulationGenetics
Average 90 stars, based on 1 article reviews
trinity short read assembler - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

90
Oxford Nanopore canu assembly of hifi reads
Left: A) Two hypothetical reads are shown with sequencing errors highlighted in red. B) The first step of HiCanu is to compress homopolymers, which obscures homopolymer length errors but retains enough information to accurately distinguish reads from different genomic loci. C) Overlaps are then computed for the compressed reads, and remaining errors are identified by examining the alignment pileups (gray rectangle). D) Finally, after correcting the identified errors (blue) and ignoring indels in regions of known systematic error (gray), the resulting overlap is 100% identical. Right: Sequence identity of reads from a 20 kbp <t>HiFi</t> library measured against <t>the</t> <t>CHM13</t> chromosome X reference sequence v0.7 after each step of HiCanu processing (Supplementary Note 1). Separate boxplots are shown for raw HiFi reads (init), homopolymer-compressed reads (compressed), OEA-corrected reads (corrected), and corrected reads after ignoring differences in microsatellite repeats (masked). The median read identity, indicated by solid segments, increases from less than 99.9% to 100% (note that the plots show an y-range of 99.65–100%). Supplementary Table 1 also shows how HiCanu processing increases the percentage of perfectly-aligned (100% identity) HiFi reads from less than 1% to over 97%.
Canu Assembly Of Hifi Reads, supplied by Oxford Nanopore, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/assembled+reads/bio_rxiv__2020__03__14__992248-283-35-41?v=Oxford+Nanopore
Average 90 stars, based on 1 article reviews
canu assembly of hifi reads - by Bioz Stars, 2026-08
90/100 stars
  Buy from Supplier

Image Search Results


(a) Read based detection of Patescibacterium . A total of 580 SRA metagenomes, out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .

Journal: bioRxiv

Article Title: Proposal of Patescibacterium danicum gen. nov., sp. nov. in the ubiquitous ultrasmall bacterial phylum Patescibacteriota phyl. nov.

doi: 10.1101/2024.11.01.618662

Figure Lengend Snippet: (a) Read based detection of Patescibacterium . A total of 580 SRA metagenomes, out of 248,559, contained hits to Patescibacterium , 480 of which had associated latitude/ longitude metadata and are shown here. Circle diameter indicates the number of samples per location cluster, and darker colors represent higher relative abundances (see legend). For display purposes, the abundance was capt at 1%. (b) Most common habitat types of Patescibacterium among the 580 SRA metagenomes. The list is based on the NCBI “organism” field, associated with NCBI BioSamples of metagenomic data, and has been manually curated to combine overlapping habitats. The values are counts of metagenomes per habitat. The original table is provided as Table S15 .

Article Snippet: The closed metagenome assembled genome (MAG) ‘ABY1’ was obtained via differential coverage binning of Oxford Nanopore Technology (ONT) long read assemblies polished by Illumina short read data from Mariagerfjord wastewater treatment plant (WWTP) (SAMN14825711) and published previously ( ).

Techniques:

Descriptive characteristics for three draft taro genome assemblies. The pseudochromosome-level taro assembly (“Ps_chr”) was composed from a linked-read assembly (“LR”) that was gap-filled using  contigs  assembled from nanopore MinION long-reads (merge step, “Merged”), filtered for assembly artifacts, and then concatenated into pseudochromosomes using a linkage map. Kilobase = kb

Journal: G3: Genes|Genomes|Genetics

Article Title: Taro Genome Assembly and Linkage Map Reveal QTLs for Resistance to Taro Leaf Blight

doi: 10.1534/g3.120.401367

Figure Lengend Snippet: Descriptive characteristics for three draft taro genome assemblies. The pseudochromosome-level taro assembly (“Ps_chr”) was composed from a linked-read assembly (“LR”) that was gap-filled using contigs assembled from nanopore MinION long-reads (merge step, “Merged”), filtered for assembly artifacts, and then concatenated into pseudochromosomes using a linkage map. Kilobase = kb

Article Snippet: We sequenced and assembled a taro genome using a linked-read sequencing strategy, with genome contiguity improved through gap filling and scaffolding using contigs assembled from Oxford Nanopore MinIon long-reads and linkage map results from a mapping population for TLB-resistance.

Techniques:

Repetitive content of taro ( Colocasia esculenta ) and great duckweed ( Spirodela polyrhiza ) genome assembles. Total repeat content was quantified using de novo repeat libraries constructed with RepeatModeler and screened with RepeatMasker. The percent (%) of sequence is relative to each individual assembly’s total length excluding runs of NNN”s between scaffolded  contigs.  Short and long interspersed elements are denoted as SINEs and LINEs

Journal: G3: Genes|Genomes|Genetics

Article Title: Taro Genome Assembly and Linkage Map Reveal QTLs for Resistance to Taro Leaf Blight

doi: 10.1534/g3.120.401367

Figure Lengend Snippet: Repetitive content of taro ( Colocasia esculenta ) and great duckweed ( Spirodela polyrhiza ) genome assembles. Total repeat content was quantified using de novo repeat libraries constructed with RepeatModeler and screened with RepeatMasker. The percent (%) of sequence is relative to each individual assembly’s total length excluding runs of NNN”s between scaffolded contigs. Short and long interspersed elements are denoted as SINEs and LINEs

Article Snippet: We sequenced and assembled a taro genome using a linked-read sequencing strategy, with genome contiguity improved through gap filling and scaffolding using contigs assembled from Oxford Nanopore MinIon long-reads and linkage map results from a mapping population for TLB-resistance.

Techniques: Construct, Sequencing

Evaluating the pros and cons of selected sequencing and mapping technologies.

Journal: Current opinion in insect science

Article Title: Recent Advances and Future Perspectives in Vector-omics

doi: 10.1016/j.cois.2020.05.006

Figure Lengend Snippet: Evaluating the pros and cons of selected sequencing and mapping technologies.

Article Snippet: Oxford Nanopore Technologies , Long read sequencing technology used for de novo genome assembly, isoform identification, and detecting epigenetic modifications. , Sequencing of long DNA molecules directly enable detection of epigenetic modifications. Can sequence transcripts for full cDNA and direct RNA sequencing for nucleotide modification detection. Length of reads are only limited by the DNA isolation and sequencing library. Sequencing devices are portable and scalable. Real time analysis is possible for rapid workflows. Several insect tissues have been sequenced showing promise for the vector field. Continuing community-guided improvements to chemistry , High error rate but can be corrected by overlap consensus. Sensitive to long homopolymers. Requires hands-on training to operate. Limited protocols for vectors. Constant updates to kits and user interface which make it challenging for comparison and to keep pace. , ~15 kb/2Mb* (in theory, read length is only limited by the DNA input).

Techniques: Sequencing, Plasmid Preparation, RNA Sequencing Assay, Modification, DNA Extraction, Comparison, Scaffolding, Variant Assay, Produced, Immunoprecipitation, Genome Wide, Expressing, Binding Assay, Activity Assay, Amplification

Left: A) Two hypothetical reads are shown with sequencing errors highlighted in red. B) The first step of HiCanu is to compress homopolymers, which obscures homopolymer length errors but retains enough information to accurately distinguish reads from different genomic loci. C) Overlaps are then computed for the compressed reads, and remaining errors are identified by examining the alignment pileups (gray rectangle). D) Finally, after correcting the identified errors (blue) and ignoring indels in regions of known systematic error (gray), the resulting overlap is 100% identical. Right: Sequence identity of reads from a 20 kbp HiFi library measured against the CHM13 chromosome X reference sequence v0.7 after each step of HiCanu processing (Supplementary Note 1). Separate boxplots are shown for raw HiFi reads (init), homopolymer-compressed reads (compressed), OEA-corrected reads (corrected), and corrected reads after ignoring differences in microsatellite repeats (masked). The median read identity, indicated by solid segments, increases from less than 99.9% to 100% (note that the plots show an y-range of 99.65–100%). Supplementary Table 1 also shows how HiCanu processing increases the percentage of perfectly-aligned (100% identity) HiFi reads from less than 1% to over 97%.

Journal: bioRxiv

Article Title: HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads

doi: 10.1101/2020.03.14.992248

Figure Lengend Snippet: Left: A) Two hypothetical reads are shown with sequencing errors highlighted in red. B) The first step of HiCanu is to compress homopolymers, which obscures homopolymer length errors but retains enough information to accurately distinguish reads from different genomic loci. C) Overlaps are then computed for the compressed reads, and remaining errors are identified by examining the alignment pileups (gray rectangle). D) Finally, after correcting the identified errors (blue) and ignoring indels in regions of known systematic error (gray), the resulting overlap is 100% identical. Right: Sequence identity of reads from a 20 kbp HiFi library measured against the CHM13 chromosome X reference sequence v0.7 after each step of HiCanu processing (Supplementary Note 1). Separate boxplots are shown for raw HiFi reads (init), homopolymer-compressed reads (compressed), OEA-corrected reads (corrected), and corrected reads after ignoring differences in microsatellite repeats (masked). The median read identity, indicated by solid segments, increases from less than 99.9% to 100% (note that the plots show an y-range of 99.65–100%). Supplementary Table 1 also shows how HiCanu processing increases the percentage of perfectly-aligned (100% identity) HiFi reads from less than 1% to over 97%.

Article Snippet: This included Oxford Nanopore UL Canu assemblies presented by ( ) for HG0002 (80x Guppy HAC 2.3.5) and HG00733 (50x Guppy HAC 2.3.5); Canu + Racon assembly presented by ( ); HG002 Canu assembly of HiFi reads presented by ( ); Oxford Nanopore Canu assembly for CHM13 (40x + 80x UL Guppy HAC 3.1.5) presented by ( ); HiFi + Hi-C assemblies for HG002 presented by ( ); HiFi + Strand-seq assemblies for HG0733 presented by ( ).

Techniques: Sequencing

HiCanu assembly of the 20 kbp HiFi dataset (left) and Canu assembly of an ultra-long Nanopore dataset (right). White regions indicate gaps in the current reference genome, while each gray and black block indicates a continuous contig alignment. Color switches from gray to black represent either the end of a contig or an alignment break. Assemblies were aligned to GRCh38 using MashMap and plots were generated using coloredChromosomes as previously described ( ; ). Note that some chromosomes (e.g. chrX) are better resolved by the Nanopore assembly due to the presence of nearperfect repeats. At the same time, chromosomes containing more diverged repeats (e.g. chr7 and chr16) are better resolved by the HiFi assembly. We note that some gaps in the HiFi assembly are caused by sequence-specific biases of current HiFi sequencing protocols (Supplementary Note 4). The red box highlights the β-defensin gene cluster on chromosome 8 which is split in both assemblies and detailed in .

Journal: bioRxiv

Article Title: HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads

doi: 10.1101/2020.03.14.992248

Figure Lengend Snippet: HiCanu assembly of the 20 kbp HiFi dataset (left) and Canu assembly of an ultra-long Nanopore dataset (right). White regions indicate gaps in the current reference genome, while each gray and black block indicates a continuous contig alignment. Color switches from gray to black represent either the end of a contig or an alignment break. Assemblies were aligned to GRCh38 using MashMap and plots were generated using coloredChromosomes as previously described ( ; ). Note that some chromosomes (e.g. chrX) are better resolved by the Nanopore assembly due to the presence of nearperfect repeats. At the same time, chromosomes containing more diverged repeats (e.g. chr7 and chr16) are better resolved by the HiFi assembly. We note that some gaps in the HiFi assembly are caused by sequence-specific biases of current HiFi sequencing protocols (Supplementary Note 4). The red box highlights the β-defensin gene cluster on chromosome 8 which is split in both assemblies and detailed in .

Article Snippet: This included Oxford Nanopore UL Canu assemblies presented by ( ) for HG0002 (80x Guppy HAC 2.3.5) and HG00733 (50x Guppy HAC 2.3.5); Canu + Racon assembly presented by ( ); HG002 Canu assembly of HiFi reads presented by ( ); Oxford Nanopore Canu assembly for CHM13 (40x + 80x UL Guppy HAC 3.1.5) presented by ( ); HiFi + Hi-C assemblies for HG002 presented by ( ); HiFi + Strand-seq assemblies for HG0733 presented by ( ).

Techniques: Blocking Assay, Generated, Sequencing

RepeatMasker of tig00006497 reveals three α-satellite HOR arrays that reside within the chromosome 19 centromere (D19Z1, D19Z2?, and D19Z3; marked with black bars). These HOR arrays are 606 kbp, 289 kbp, and 3.96 Mbp in length, respectively, and are composed of a 13-mer, a complex higher-order HOR, and a dimeric HOR unit, respectively. The HOR repeat underlying D19Z2 shares limited sequence identity with the pG-A16 repeat previously described ( ; ; ) and, therefore, is designated with a question mark. The α-satellite HOR arrays have relatively uniform coverage of HiFi and ultra-long Oxford Nanopore data, except for a drop in Oxford Nanopore sequencing coverage over the D19Z1 array, which may be due to a mis-assembly, read mis-mapping, or biases in sequencing. The HiFi coverage plot shows fold coverage of the most common base (black) and the second most common base (red).

Journal: bioRxiv

Article Title: HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads

doi: 10.1101/2020.03.14.992248

Figure Lengend Snippet: RepeatMasker of tig00006497 reveals three α-satellite HOR arrays that reside within the chromosome 19 centromere (D19Z1, D19Z2?, and D19Z3; marked with black bars). These HOR arrays are 606 kbp, 289 kbp, and 3.96 Mbp in length, respectively, and are composed of a 13-mer, a complex higher-order HOR, and a dimeric HOR unit, respectively. The HOR repeat underlying D19Z2 shares limited sequence identity with the pG-A16 repeat previously described ( ; ; ) and, therefore, is designated with a question mark. The α-satellite HOR arrays have relatively uniform coverage of HiFi and ultra-long Oxford Nanopore data, except for a drop in Oxford Nanopore sequencing coverage over the D19Z1 array, which may be due to a mis-assembly, read mis-mapping, or biases in sequencing. The HiFi coverage plot shows fold coverage of the most common base (black) and the second most common base (red).

Article Snippet: This included Oxford Nanopore UL Canu assemblies presented by ( ) for HG0002 (80x Guppy HAC 2.3.5) and HG00733 (50x Guppy HAC 2.3.5); Canu + Racon assembly presented by ( ); HG002 Canu assembly of HiFi reads presented by ( ); Oxford Nanopore Canu assembly for CHM13 (40x + 80x UL Guppy HAC 3.1.5) presented by ( ); HiFi + Hi-C assemblies for HG002 presented by ( ); HiFi + Strand-seq assemblies for HG0733 presented by ( ).

Techniques: Sequencing, Nanopore Sequencing

Top: Nucmer self-alignment dot plots of the CHM13 reference defensin region at different alignment stringencies (Methods). A: greater than 7 kbp repeats at 98% identity. B: greater than 7 kbp repeats at 99.9% identity. Purple/blue indicates same/reverse strand matches. C: Icarus visualization of contig alignments from both HiFi-based (Canu, HiCanu, Peregrine) and ultra-long Nanopore-based assemblies (Canu ONT and Flye ONT ) produced by QUAST . White space in the alignment figure indicates the assembly was fragmented into short contigs (<50 kbp). Red color indicates mis-assembled contigs. The HiCanu assembly breaks at two of three segmental duplication instances which share high sequence similarity (black arrows) and at a region of systematic HiFi coverage depletion (red arrow).

Journal: bioRxiv

Article Title: HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads

doi: 10.1101/2020.03.14.992248

Figure Lengend Snippet: Top: Nucmer self-alignment dot plots of the CHM13 reference defensin region at different alignment stringencies (Methods). A: greater than 7 kbp repeats at 98% identity. B: greater than 7 kbp repeats at 99.9% identity. Purple/blue indicates same/reverse strand matches. C: Icarus visualization of contig alignments from both HiFi-based (Canu, HiCanu, Peregrine) and ultra-long Nanopore-based assemblies (Canu ONT and Flye ONT ) produced by QUAST . White space in the alignment figure indicates the assembly was fragmented into short contigs (<50 kbp). Red color indicates mis-assembled contigs. The HiCanu assembly breaks at two of three segmental duplication instances which share high sequence similarity (black arrows) and at a region of systematic HiFi coverage depletion (red arrow).

Article Snippet: This included Oxford Nanopore UL Canu assemblies presented by ( ) for HG0002 (80x Guppy HAC 2.3.5) and HG00733 (50x Guppy HAC 2.3.5); Canu + Racon assembly presented by ( ); HG002 Canu assembly of HiFi reads presented by ( ); Oxford Nanopore Canu assembly for CHM13 (40x + 80x UL Guppy HAC 3.1.5) presented by ( ); HiFi + Hi-C assemblies for HG002 presented by ( ); HiFi + Strand-seq assemblies for HG0733 presented by ( ).

Techniques: Produced, Sequencing